今天進入後端語音模組與 API 串接階段,完成項目包含:
spoken_summary 白話文轉化為 .mp3 語音檔。app.py) — 新增 /analyze-prescription API 端點,支援接收藥袋圖片上傳、執行 VLM 解析並回傳 JSON 與語音檔檔名。/static/audio/<filename>) — 提供前端(LINE Bot / Web 儀表板)直接存取與播放語音檔的 URL。許多使用者裝置(如瀏覽器或手機)內建的 TTS 引擎發音品質差異大,常出現發音生硬、斷句不自然或缺少繁體中文語音包的問題。
將語音合成交由 Python 後端處理有幾個優勢。採用 Google gTTS 雲端語音確保發音品質一致,長輩能聽得清楚。前端手機只需下載 .mp3 並播放,不需要額外下載或初始化龐大的語音包資源。
app.py)開啟專案根目錄下的 app.py,更新為以下完整程式碼:
import os
import json
import uuid
from flask import Flask, request, jsonify, send_from_directory
from dotenv import load_dotenv
from google import genai
from google.genai import types
from PIL import Image
from gtts import gTTS
# 1. 載入環境變數與初始化
load_dotenv()
api_key = os.getenv("GEMINI_API_KEY")
if not api_key:
raise ValueError("❌ 錯誤:找不到 GEMINI_API_KEY,請檢查 .env 設定!")
client = genai.Client(api_key=api_key)
app = Flask(__name__)
# 設定語音檔儲存目錄
AUDIO_DIR = os.path.join(os.getcwd(), 'static', 'audio')
os.makedirs(AUDIO_DIR, exist_ok=True)
# 2. 定義 JSON Schema
prescription_schema = {
"type": "OBJECT",
"properties": {
"spoken_summary": {
"type": "STRING",
"description": "適合唸給長輩聽的白話文總結,語氣溫柔親切。"
},
"medicines": {
"type": "ARRAY",
"items": {
"type": "OBJECT",
"properties": {
"name": {"type": "STRING", "description": "藥品中文名稱或商品名"},
"type": {"type": "STRING", "description": "口服或外用"},
"frequency": {"type": "STRING", "description": "服用頻率,如:每日二次"},
"dosage": {"type": "STRING", "description": "每次劑量,如:1顆、0.5顆"},
"timing": {"type": "STRING", "description": "吃藥時間點,如:飯後、睡前"},
"warning": {"type": "STRING", "description": "重要警語或注意事項,無則填空字串"}
},
"required": ["name", "type", "frequency", "dosage", "timing"]
}
}
},
"required": ["spoken_summary", "medicines"]
}
@app.route('/ping', methods=['GET'])
def ping():
return jsonify({"status": "online", "service": "PrescriptionVLM Engine"}), 200
@app.route('/analyze-prescription', methods=['POST'])
def analyze_prescription():
# 檢查是否有傳入圖片檔案
if 'image' not in request.files:
return jsonify({"error": "未提供圖片檔案 (key 名稱應為 'image')"}), 400
file = request.files['image']
if file.filename == '':
return jsonify({"error": "未選擇上傳的圖片"}), 400
try:
# 1. 讀取影像
image = Image.open(file.stream)
# 2. 呼叫 Gemini VLM 模型
prompt = """
你是一位專業且細心的藥師助手。請分析這張藥袋照片:
1. 將藥品分類為口服或外用,並精準提取名稱、頻率、劑量與吃藥時間。
2. 針對高齡長輩,撰寫一段溫柔白話的 spoken_summary,適合 TTS 語音播報。
3. 若有重要注意事項,請填入 warning 欄位。
"""
config = types.GenerateContentConfig(
response_mime_type="application/json",
response_schema=prescription_schema
)
response = client.models.generate_content(
model='gemini-3.6-flash',
contents=[image, prompt],
config=config
)
result_data = json.loads(response.text)
# 3. 將 spoken_summary 轉為 mp3 語音檔
spoken_text = result_data.get("spoken_summary", "解析完成。")
filename = f"speech_{uuid.uuid4().hex[:8]}.mp3"
filepath = os.path.join(AUDIO_DIR, filename)
tts = gTTS(text=spoken_text, lang='zh-tw')
tts.save(filepath)
# 4. 組裝回傳 JSON 格式
result_data['audio_url'] = f"/static/audio/{filename}"
return jsonify({
"status": "success",
"data": result_data
}), 200
except Exception as e:
return jsonify({"error": f"伺服器處理失敗: {str(e)}"}), 500
@app.route('/static/audio/<filename>', methods=['GET'])
def get_audio(filename):
"""提供語音檔下載/播放靜態路由"""
return send_from_directory(AUDIO_DIR, filename)
if __name__ == '__main__':
app.run(host='0.0.0.0', port=5000, debug=True)
1. 啟動 Flask 後端服務
於 Terminal 執行,輸入(Ctrl+C 可取消執行 app.py):
python app.py
2. 使用 cURL 或 Postman 測試 API 請求
開啟新的 Terminal(終端機)分頁,執行以下 cURL 指令上傳測試藥袋圖片:
curl -X POST http://127.0.0.1:5000/analyze-prescription \
-F "image=@test_rx.jpg"
伺服器回傳包含 status、data(藥品陣列與白話文)以及 audio_url(如 /static/audio/speech_a1b2c3d4.mp3)的 JSON 資料。
http://127.0.0.1:5000/static/audio/speech_dbefbd4a.mp3
(請替換為實際回傳的檔名),即可聆聽繁體中文藥袋說明語音。
測試成功後,將 Day 06 的程式碼提交至 GitHub:
git add .
git commit -m "保留雙引號 改填寫自己要記錄的標記 ex.鐵人賽第六天"
git push
今天將語音合成模組(gTTS)整合至 Flask API 中,實現了藥袋照片轉語音的後端服務。傳入照片、Gemini 解析、自動生成 .mp3 檔案這些步驟一起串聯。
明天(Day 07)進行 SQLite 資料庫整合與用藥歷史 API 實作,把解析出的藥品資料與語音路徑存入資料庫,建立完整的歷史紀錄管理機制。